Papers with corpus creation methodology
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
MuST-C: a Multilingual Speech Translation Corpus (N19-1)
Copied to clipboard
| Challenge: | Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora. |
| Approach: | They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages. |
| Outcome: | The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages. |